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Abstract 

Wc describe two eonstructions of (very) dense graphs which are edge disjoint unions of large 
induced matchings. The first construction exhibits graphs on N vertices with (^) —o{N'^) edges, 
which can be decomposed into pairwise disjoint induced matchings, each of size 
second construction provides a covering of all edges of the complete graph Kn by two graphs, 
each being the edge disjoint union of at most N^^^ induced matchings, where 6 > 0.058. 
This disproves (in a strong form) a conjecture of Meshulam, substantially improves a result of 
Birk, Linial and Meshulam on communicating over a shared channel, and (slightly) extends the 
analysis of Hastad and Wigderson of the graph test of Samorodnitsky and Trevisan for linearity. 
Additionally, our constructions settle a combinatorial question of Vempala regarding a candidate 
rounding scheme for the directed Steiner tree problem. 
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1 Introduction 



1.1 B ackgr ound 

Dense graphs consisting of large pairwise edge disjoint induced matchings have found several ap- 
plications in Combinatorics, Complexity Theory and Information Theory. Call a graph G = {V, E) 
an (r, i)-Ruzsa-Szemeredi graph ((r, t)-RS graph, for short) if its set of edges consists of t pairwise 
disjoint induced matchings, each of size r. The total number of edges of such a graph is clearly rt. 
Graphs of this type are useful when both r and t are relatively large as a function of the number of 
vertices A^. There are several known interesting constructions, relying on a variety of techniques. 

The first surprising construction was given by Ruzsa and Szemeredi [21], who applied a result of 
Behrend [7] about the existence of dense subsets of {1, 2, 0(A'^)} containing no 3-term arithmetic 
progressions to prove that there are (r, t)-RS graphs on N vertices with r = m^^g^) ^ ~ -^/3- 
They applied this construction, together with the regularity lemma of Szemeredi [24], to settle an 
extremal problem of Brown, Erdos and Sos [s, 9], showing that the maximum possible number of 
edges in a 3-uniform hypergraph on N vertices which contains no 6 vertices spanning at least 3 
edges is bigger than A^^~^ and smaller than eA^^, for any e > 0, provided N > iVo(e). See also 
[13], [5] for more details about this problem, its extensions, and their connection to (r, t)-RS graphs. 

Note that the above construction provides graphs with vertices and ^owiogN) edges, that is, 
rather dense graphs, but still ones in which the number of edges is o{N'^). These graphs and some 
appropriate variants have been used by the first author in [1], to show that the problem of testing 
i^-freeness in graphs requires a super-polynomial (in 1/e) number of queries if and only if H is not 
bipartite. The proof for one sided error algorithms is given in [1], and an extension for two sided 
algorithms is described in [.'>]. A similar application of these graphs for testing induced i?-freeness 
appears in [4], and yet another very recent application showing that testing graph-perfectness 
requires a super polynomial number of queries appears in [2]. 

The above graphs have also been applied by Hastad and Wigderson []!)] to give an improved 
analysis of the graph test of Samorodnitsky and Trevisan for linearity and for PCP with low 
amortized complexity [22]. 

Another construction of (r, f)-RS graphs on A^ vertices, with r = N/3—o{N) and t = A^^(i/^°siog^) 
was given by Fischer et. al. in [14]. Note that the matchings here are of linear size, but their num- 
ber is much smaller than in the original construction of Ruzsa and Szemeredi. The construction 
here is combinatorial, and Fischer et. al. use these graphs to establish an A^^(i/'°si°s^) lower 
bound for testing monotonicity in general posets. 

Yet another construction was obtained by Birk, Linial and Meshulam [10], and in an improved 
form by Meshulam [20]. For the application in [10] it is crucial to obtain graphs with positive density. 
Indeed, the graphs here are (r, t)-graphs on A^ vertices with r = (log A^)^(^°siog^/(iogiogiog^) ) 
and t roughly Thus, their number of edges is about N^/2A. The method here relies on 

a construction of low degree representation of the OR function, due to Barrington, Beigel and 
Rudich [6]. The application in [10] is in Information Theory, the graphs are applied to design an 
efficient deterministic scheduling scheme for communicating over a shared directional multichannel. 

Interestingly, none of these constructions address the question of whether or not an (r, f)-RS 
graph can simultaneously have positive density and yet be an edge disjoint union of polynomially 
large induced matchings. This range of parameters is important for some applications - especially 
ones in which there is a tradeoff between the number of missing edges and the number of induced 
matchings needed to cover the graph. Indeed, Meshulam [20] conjecture that there are no such 
graphs. We are able to disprove this conjecture in the strongest possible sense: The density of our 
construction is 1 — o(l) and yet r is nearly linear in A^. We also give a number of applications of 
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our constructions. 



1.2 Our Results 

We construct (r, t)-RS graphs on A'" vertices with rt = {1 — o(l))(^), and r = N^~°^^h Thus, 
not only can we have graphs with positive edge density which are edge disjoint union of induced 



Kn into two subgraphs, each being a union of at most N"^^^ induced matchings, where 6 > 0.058. 
The main difference between the new constructions presented here and the previous ones mentioned 
above, is that the graphs constructed here are of density 1 — o(l), that is, almost all edges of the 
complete graph Kn are covered, and yet all these edges can be partitioned into large pairwise 
disjoint induced matchings. This surprising property turns out to be useful in various applications. 

Our first construction is geometric, and is inspired by the recent work of Fox and Loh [16] on 
dense graphs in which every edge is contained in at least one triangle and yet no edge is contained in 
too many triangles. The construction follows the basic approach of Fox and Loh (slightly modified 
according to the remark of the first author, mentioned at the end of [16]) with different parameters. 
An additional (simple) argument is required in decomposing sparse graphs into not too many 
induced matchings. Our second construction applies some basic tools from Coding Theory. Also 
we make us of the regularity lemma and some combinatorial and entropy based techniques to prove 
lower bounds for these questions. 

It is worth noting that a general result of Frankl and Fiiredi [17], implies that for any fixed r, 
there are (r, t)-RS graphs G on N vertices with rt = (1 — o(l))(^). This is proved by choosing 
the non-edges of G randomly and by applying the nibble technique to obtain the existence of the 
desired matchings. Using this method, however, the induced matchings obtained are of constant 
size, whereas we are interested, crucially, in large matchings. The techniques of [17] cannot provide 
induced matchings of size exceeding 0(log A^). 

We apply our results to significantly improve the application in [10]. As mentioned earlier, Birk, 
Linial and Meshulam construct (r,t)-graphs on N vertices with r = (bg iV)f^(iogiog^/(iogiosiog^)^) 
and rt roughly The authors then use these graphs to design a communication protocol over a 
shared directional multi-channel - it is critical for the application that these graphs have positive 
density. The communication protocol based on these graphs achieves a round complexity of O(^) 
and this is a slightly better than poly-logarithmic improvement over the naive protocol for bus-based 
architectures. 

We can use our construction to achieve a round complexity of 0(N^~^) over a shared directional 
multi-channel. This is the first such protocol that is a polynomial improvement over the naive 
protocol. We can accomplish this using just two receivers per station (this corresponds to a partition 
of the edges of a complete bipartite graph into two graphs that can be decomposed into large 
induced matchings). In case we are allowed C = C(e) receivers per station, we can achieve a round 
complexity that is 0{N^~^'^) for any e > 0. Hence, previous protocols required nearly a quadratic 
number of rounds, and our protocols require only a nearly-linear number of rounds. 

Our constructions also disprove the recent conjecture of Meshulam. Moreover, we can achieve 
a density approaching one while simultaneously being able to decompose the graph into at most a 
nearly-linear number of induced matchings. 

Besides their applications to the problems of [JO] and [20], our constructions can be plugged in 
the result of [19], extending it to a new range of the parameters that may be of interest. Lastly, we 
also answer a question of Vempala [25], showing that a certain rounding scheme for the directed 
Steiner tree problem is not effective. 
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The rest of this paper is organized as fohows. In the next section we describe the two new 
constructions. Lower bounds showing that these are not far from being tight are given in Section 
3. In Section 4 we describe the apphcations of these graphs. The final Section 5 contains some 
concluding remarks and open problems. 

2 Constructions 

2.1 A Geometric Construction 

Here we construct nearly-complete graphs that can be covered by an almost linear number of 
induced matchings. These graphs will be based on a geometric construction inspired by a recent 
construction of Fox and Loh [ I (>] . 

We first describe a graph G = (V, E) and then prove that it can be slightly modified to yield 
a nearly complete (r, t)-RS graph. Set V = [C]*^ for some constant C to be chosen later. Let 
N = be the number of vertices. Each vertex x £ V will be interpreted as an integer vector in 
n dimensions with coordinates in [C] = {1,2,...,C}, where for technical reasons it is convenient 
to assume that n is even. Let n = Ex^y[\\x — y\\'2\ where x and y are sampled uniformly at random 
from V. We could compute fi, but we will not need this exact value. 

Next, we describe the set E of edges. A pair of vertices x and y are adjacent if and only if 



This condition implies that the number of missing edges is small, by a standard application of 
Hoeffding's Inequality: 



Proof: The quantity \\x — = Y17=ii-'^'i ~ Vif' hence is the sum of independent random 
variables (when x and y are chosen uniformly at random from V). Each variable is bounded in the 
range [0, C^] and hence we can apply Hoeffding's Inequality to obtain 



and this implies the Claim. ■ 

As a first step, we will cover all the edges of G by a linear number of induced subgraphs of 
small (but super-constant) maximum degree. We will then use this covering to obtain a covering 
via an almost linear number of induced matchings. Next, we describe the (preliminary) induced 
subgraphs that we use to cover G. 

We will define one subgraph Gz for each z . Let Vz (the vertex set of Gz) be 



The subgraph Gz is the induced graph on Vz- First, we prove that these subgraphs Gz do indeed 
cover the edges of G: 

Lemma 2.2. Let n > 2G . For all {x,y) G E, there is a z such that x,y gVz 
Proof: First we establish a simple Claim that will help us choose an appropriate z: 



^ - yWl - fJ- <n. 



Claim 2.1. (^) - \E\ < (^)2e-/2^* 



Pr[\ \\x — — ^\ > n] < 2e 



n/2C* 
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Claim 2.3. Let a be a vector in which the absolute value of each entry is at most C . Then there 
is a vector w where each entry is ±1/2 such that \[a,w)\ = \ X^^Li Oj^il ^ C /2 < n/A 

Proof: We can prove this by induction by considering the partial sum aiWi which we will 

assume is at most (7/2 in absolute value. We can choose Wr+i so that ar+iWr+i has the opposite 
sign of this partial sum and this implies that the partial sum X^^^i OiW^j is also at most C/2 in 
absolute value (although the sign may have changed). This completes the proof of the claim. ■ 

Let a be a vector defined as follows: if yi — Xi is even, set Oj = and otherwise set ai = yi — Xi. 
We can apply Claim 2.3 to a, but furthermore change the values of w to be zero on indices on 
which a is zero. Then |(a, w)\ is still at most C/2, and vui is ±1/2 whenever Oj is non-zero, and zero 
whenever a is zero. Set z = ± vj. Note that z ^ V because whenever is not an integer, 
must be non-zero and hence vji is ±1/2 and alternatively whenever is an integer, at is zero 

and hence Wi is zero. Consider the quantity 

\\z - x\\l = II ^-y^ +w\\2 = \\\y- ^Wl + {y-x,w) + \\w\\l 

Since x and y are adjacent in G, we have that | jHy — 2;||2~m/4 | < n/4:. Also, 111^112 < n/4. Finally, 
{y — x,w) = (a, w) since w is zero iff a is zero. Hence | \\z — j;||| — /i/4 | < n/4 + C/2 + n/4 < 3n/4, 
and an identical argument holds for bounding | ||^ — yUl — /^/4 |. Thus x,y £ Vz and the edge {x, y) 
is covered by some induced subgraph Gz- ■ 

Next we establish that the maximum degree of any induced subgraph Gz is not too large: 
Lemma 2.4. For all z £V, the maximum degree of Gz is at most (10.5)". 

Proof: Let x ^Vz and consider any neighbor y oi x that satisfies y G 14 ■ We first establish that x 
and y are close to being antipodal in the ball centered at z, and hence we can bound the number 
of neighbors of x in Vz by bounding the number of points of V in some small spherical cap around 
the antipodal point to x. 

Define the antipodal point x' = 2z — x, this is the antipodal to x with respect to the ball 
centered at z. 

Consider the parallelogram (x, z,y,x + y — z). By the Parallelogram Law, the sum of the squares 
of the four side lengths equals the sum of the squares of the lengths of the two diagonals. Therefore 
we obtain 

Ik - 2/II2 + \\x + y - 2z\\l = 2\\x - z\\l +2\\y - z\\l 
Using the definition of x' , this gives 

11^; - y\\l + \\y - x'\\l = 2\\x - z\\l + 2\\y - z\\l 

Hence ||y — x'lH = 2||x — ± 2||y — zUl — ||x — y\\% and as both ||a; — zUl and ||y — are 
approximately /i/4 (since 2;, y G Vz) and ||x — y||| is approximately /i because x is adjacent to y in 
G, this implies that ||y — x'||2 < 4n. Therefore we can bound the degree of x in Gz by the number 
lattice points in a ball of radius 2-y/n (centered at the lattice point x' .) The unit n-dimensional 
cubes centered at these lattice points are pairwise disjoint, each has volume 1, and they are all 
contained in a ball of radius 2^/n + 0.5-y/n = 2.5-y/n. Therefore, the number of these points does 
not exceed the volume of an n dimensional ball of radius 2.f)y/n. Since n is even, the volume of 
this ball is 
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where here we have used the fact that 6! > (b/e)^ for any positive integer b. This completes the 
proof of the lemma. ■ 

Lemma 2.5. Let H be a graph with maximum degree d. Then H can be covered by 0{d'^) induced 
matchings. 

Proof: Call two edges ei, 62 of H in conflict if they either share a common end or there is an edge 
in H connecting an endpoint of ei to an endpoint of 62- It is clear that any edge e of H can be in 
conflict with at most 2d - 2 + {2d - 2){d - 1) < 2d'^ other edges of H. 

Thus we can initialize each member of a set of induced matchings Mi to be the empty set, 
and for each edge e of H in its turn, add e to the first Mj for which e is not in conflict with any 
edge currently in Mj. Since e is in conflict with less than 2d'^ edges, it can added to some Mj. We 
can continue this procedure, obtaining a set of less than 2d'^ induced matchings covering all edges 
ofH. U 

It follows that we can decompose the edges of each induced subgraph Gz into at most 0{d^) < 
O((10.5)^") induced matchings, and this yields a decomposition of G into 0{Nd'^) induced match- 
ings. These matchings can additionally be made edge-disjoint, since, if any edge is multiply covered 
we can remove it from all but one of the induced matchings (and the result is still an induced match- 
ing). We have thus proved the following. 

Theorem 2.6. For every n,G with n > 2C , n even, there is a graph G on N = C"" vertices that 
is missing at most edges, for 

and can be covered by N-^ disjoint induced matchings, where 

^ ^ 21nl0.5 



Hence, for any e > we can construct a graph G on N vertices missing at most N'^ ^ edges for 
S = 5{e) = e-°(i/^), that can be covered by A^^"*""^ pairwise disjoint induced matchings. Note that 
the number of matchings is nearly linear. Note also that by splitting each of these matchings M 
into [|M|/rJ pairwise disjoint matchings, each of size exactly r = N^^*"^^ , omitting the remaining 
\M\ — r[\M\/r\ < r edges, we get an (r, t)-RS graph, where r = N^~''~^ and the number of 
missing edges is at most 2N'^~^ . As e (and hence 5 < e) can be chosen to be arbitrarily small, this 
gives, with the right choice of parameters, an (r, t)-RS graph on N vertices, with r = N'^~°^^'> and 
rt=(^)-o(iV2). 

2.2 A Construction Using Error Correcting Codes 

Here we construct nearly-complete graphs with large induced matchings using error correcting 
codes. These constructions will be incomparable to those in the previous section - the number of 
missing edges will be much smaller (in fact, the number of missing edges can be made asymptotically 
optimal as we will demonstrate in Section 3.2), but the price we pay is that the average size of an 
induced matching will only be a small power of N as opposed to nearly-linear. As we will show, 
the construction in this section will be better tailored to the application in [10] (at least for some 
values of the relevant parameters) than the construction of the previous section. 
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Throughout this section, we wih use codes over the binary alphabet as well as over a bigger 
alphabet. Let dH{x,y) be the Hamming distance between two binary strings x and y (of the 
same length). The Hamming weight of a binary string x is the number of non-zero entries - or 
equivalently the Hamming distance to the all zeros vector. We can similarly define the Hamming 
distance df{{x,y) between two vectors x and y over a larger alphabet [C] as the number of indices 
where these vectors disagree. 

Definition 2.7. A [n, k, d\ linear code C is a subspace consisting of 2^^ length 7i binary vectors such 
that for all x,y € C and x ^ y, dH{x,y) > d. We will call n the encoding length, k the dimension, 
and d the distance of the code. 

An nxk matrix A of full column rank over GF{2) is the generating matrix of a code of dimension 
k and length n consisting of all linear combinations of its columns. The distance of this code is 
exactly the minimum Hamming weight of any non-zero code word. Throughout this section we will 
make use of a particular type of code: 

Definition 2.8. Call a linear code C proper if the all ones vector is a codeword. 

It is well-known that there are linear codes that achieve the Gilbert-Varshamov Bound. In fact, 
proper codes also achieve this bound: 

Lemma 2.9. IfYli=o il) ^ 2"~^', then there is a proper [n,k,d] code. Thus, there is such a code 
in which = (1 — H{d/n))n, where H{x) = —x\og2X — (1 — a;)log2(l — x) is the binary entropy 
function. 

Proof: We can define a length n code by choosing an {n — k)xn parity check matrix as follows: for 
each of the first n — 1 columns, choose each vector uniformly at random. Choose the last column to 
be the parity of the preceding n — 1 columns. Let the matrix be B. Then the code C is defined as 
C = {a; G {0, l}'^\Bx = 0}. Clearly 1 G C by construction. Since this code is linear, the minimum 
distance is exactly the minimum Hamming weight of any non-zero codeword. This quantity is 
exactly the smallest number of columns of B that sum to the all zeros vector. 

Claim 2.10. For any fixed set S of columns of B, the probability that the sum is the all zeros 
vector is exactly 2~('"~'^) 

Proof: If this set S does not contain the last column, then the sum of the columns is distributed 
uniformly on {0, 1}""*"'. If the set 5 does contain the last column, then the sum of the columns is 
exactly the sum of the columns not in the set - i.e. [n] — S - and hence is also distributed uniformly 
on {0,1}"-^ ■ 

So the probability that any set of at most d columns sums to the all zero vector is at most 
2-{n-k) ^d^^ ^n-^^ ■£ ^d^^ (n-^ ^ 2^-^ there is a parity check matrix B so that the code C has 

distance at least d + 1 > d. This code has dimension at least k because there are n — k constraints 
imposed by the parity check matrix. If these constraints are linearly independent, then the code 
has dimension exactly k. If these constraints are not linearly independent, we can add additional 
constraints until the code has dimension exactly k and the distance of the code cannot decrease as 
we add these constraints. ■ 

Claim 2.11. Let A be the generating matrix for a proper [n, k, d] code C with d > 1. Then deleting 
any row of A results in a generating matrix A' for a proper [n — l,k,d — 1] code. 
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Proof: Note that Q' = {A'x\x G {0, l}'^} and hence C (defined by the generating matrix A') has 
dimension k, as no nontrivial linear combination of the columns of A' can be the zero vector, by 
the assumption d > 1. The all ones vector is still a codeword since Ax = 1 implies that A'x = 1. 
Finally, the minimum distance of C is the minimum Hamming weight of any non-zero code word, 
and the Hamming weight of any codeword in C decreases by at most one by deleting any index. ■ 

Throughout this section let C = C„ be an [n, k, d] code, and let Cn-i, Qn-2, ■■■Qn-d+i be proper 
[n — 1, k, d — 1], [n — 2, k, d — 2], ... and [n — d + l,k, 1] codes, respectively. 

Next, we define a graph G = {V, E) that will be the focus of this section. Let V = [C]" and set 
N = \V\ = C". Consider two vertices a,h e V where a = (ai, 02, ...a„) and h = (61, 62, ...6n) for 
Oj, 6i G [C]. There is an edge between a and h if and only if dnia, b) = ^11=1 ^at^bi > n — d. 

It is easy to count the number of missing edges. Indeed, in the complement of G each vertex a 
is connected to all vertices b so that = 5j for at least d indices i. As the number of missing edges 
is half the sum of degrees in the complement this gives: 

Claim 2.12. 

■^)-i^iac«|:0(c-ir-. 

Lemma 2.13. If ^ > then 



to 

i=d ^ ^ 



{G-Xf-' < ( ^ )C"(C- 1) 



Proof: Using the inequality (^) < {n/dy '^(^) we obtain a bound 

n ^ s ^ s n—d 

i=d^^'' ^ ^ i=o 

which implies the Lemma. ■ 

Hence the number of missing edges in G is at most for e = 1 + -^('^/")+(^^^^/^ *°g2(C 



Next, we describe the induced matchings that are used to cover the edges in G. In order to do 
so, we will define an equivalence relation over edges of G. In particular, this will be an equivalence 
relation over ordered pairs (a, 6), where a = (ai, 02, ...an) and b = (61, 62, ■■■bn)-, under the condition 
that dnia, b) > n — d. 

Definition 2.14. Let S <Z [n], l^j = r and let (a, b) be a pair of vertices in V where S = {i\ai = bi\. 
Let X be a {0, vector. Let [n] — 5 = {11,12-, ■■■in-r] and ii < 12, ... < in-r- Then the x-flip of 

(a, b) is a pair (c, d) such that for all i ^ S, Ci = ai = bi = di and for all i = ij ^ S (i.e. i is the j^^ 
smallest index not in S), Ci = ai,di = bi if Xj = and otherwise Ci = bi,di = ai. 

Informally, the n — r indices not in S are mapped in order to the n — r bits in x and the 
corresponding locations in a and b are swapped if and only if the corresponding bit of x is one. 

Definition 2.15. We wih define a pair (a, 6) ~ {a',b') iff 5 = {i\ai = h} = S' = {i\a[ = b[}, 
\S\ < d and furthermore there is an x G C such that (a', b') is the x-flip of (a, b). 
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Next we will establish that this relation ~ is indeed an equivalence relation, and that it is 
actually a relation on unordered pairs, that is (a, b) ~ (b, a) for all a, b: 



Claim 2.16. (a, 6) ~ (6,0) 

This follows because the code Qn-r is proper (for all r < d), and hence the all ones vector 1 lies 
in Qn-r aiid (6, a) is the 1-flip of (a, b). 

Claim 2.17. (a, b) ~ (c, d) iff (c, d) ~ (a, 6) 

Proof: By symmetry we only need to establish one direction. Suppose {a,b) ~ {c,d). Then 
S = {i\ai = bi} = S' = {i\ci = di}. Let (c, d) be an x-flip of (a, 6) (where x e Cn-r)- Then (a, 6) is 
also the x-flip of (c, d). ■ 

Claim 2.18. (a, 6) ~ (c, d) and (c, d) ~ (e, /) implies {a,b) ~ (e,/) 

Proof: Again note that S = {i\ai = bi} = S' = {i\ci = di} = S" = {i\ei = fi}. Let x,y G Qn-r be 
such that (c, d) is the x-flip of {a,b) and (e,/) is the y-flip of {c,d). Then x -\- y £ Cn~r since the 
code is linear, and (e, /) is the x + y-flip of (a, b). ■ 

This immediately implies: 

Lemma 2.19. The relation ~ is an equivalence relation over unordered pairs {a,b) which have 
Hamming distance > n — d. 

Since each code Cn-r (for r < d) has dimension k, each equivalence class has size exactly 2^. 

Lemma 2.20. Each equivalence class is an induced matching consisting of 2^~^ edges. 

Proof: Consider two edges (a, 6) and (e, /) in the same equivalence class. Let 5 = {i\ai = bi} = 
{i\^i = fi} where \S\ = r (< d). Let (e, /) be the x-flip of (a, 6) for x G Qn-r- Since the code Qn-r 
has distance at least d — r, the Hamming weight of x is at least d — r. Consider the Hamming 
distance between a and /. Each index z € 5" is an index at which a and / agree (i.e. Oj = fi). 
Furthermore, there is a bijection between indices in x that are set to one and indices outside of the 
set S, for which a and / agree. So the vectors a and / agree on at least r+{d — r) = d indices, and 
hence af is not an edge in G. Since (e, /) and (/, e) are in the same equivalence class the above 
argument also shows that ae,be and bf are nonedges. ■ 

If we use one induced matching for each equivalence class, then each edge in G is covered exactly 
once and hence the number of induced matchings needed to cover G is < 

Theorem 2.21. For every n,d,G such that ^ > , there is a graph G on N = C" vertices 
that is missing at most edges, for 

^ ^ ^ ^ Hid/n) + {l-d/n)log^{C-l) ^ 
loggC 

and can be covered by disjoint induced matchings, where 

logjC 
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In particular, for any e > 0, there is a graph G on N vertices missing at most ges 
that can be covered by A^^"'^'^ induced matchings. This is obtained by choosing C for which 
log2 C = e(l/e) and d/n = ^ - G(e). 

Also we can choose C = 34 and d = 0.19n, in which case e, / < 1.942. Thus we can cover the 
edges of a complete graph on 2N vertices by two graphs Gi (set to G with N replaced by 2N in the 
above construction) and G2 (set to the complement of G) , where the number of induced matchings 
needed to cover the edges of Gi is 0{N^-^) for 5 > 0.058, and the same holds for G2. For the 
applications we need that the above statement holds also for covering all edges of the complete 
bipartite graph K^^n by two such graphs G'l and G2-this clearly follows by splitting the vertices 
of K2N arbitrarily into two equal classes and by defining G[, for i = 1, 2, to be the graph obtained 
from Gi by keeping only the edges that have one endpoint in each class. 

3 Limits 

3.1 Triangle Removal Lemma 

The connection between the triangle removal lemma and the existence of (r, t)-RS graphs is well 
known since the work of Ruzsa and Szemeredi, for completeness we include the argument. 

Proposition 3.1. If there exists an {r,t)-RS graph on N vertices, then there exists a graph on 
N + 1 vertices with at least 3rt/2 edges, in which every edge is contained in exactly one triangle. 
Thus one has to delete at least rt/2 edges to destroy all triangles and yet the graph contains only 
rt/2 triangles. 

Proof: Let G be an (r, t)-RS graph on N vertices. Then its number of edges is rt and hence, by a 
well known simple result, it contains a bipartite subgraph G' = {U,V,E') with at least rt/2 edges. 
Clearly, these edges can be covered by t induced matchings Mi, M2, ...Mt, and we can assume that 
these matchings are pairwise edge disjoint. 

For each matching Mi , add an additional vertex Wi and connect Wi to the endpoints of all edges 
in Mi. The resulting graph H = {U,V,W, Eh) is tripartite, has N + t vertices and contains \E'\ 
triangles. The critical property of this construction is that each edge of is in a unique triangle. 
Indeed, there is a natural set of \E'\ triangles in H - each such triangle is specified by an edge 
{u,v) € E' and if this edge is contained in the matching Mj, this edge is mapped to the triangle 
{u,v,Wi) in H. There are in fact no other triangles in H: Let T = (a,b,c) be a triangle in H. 
Since H is tripartite, there must be exactly one vertex from each set U, V and W in the set a, 6, c. 
Suppose that a £ U and b £ V. Then let Mi be the unique matching containing the edge (a, 6). 
Suppose c = Wj ^ Wi. This implies that the matching Mj covers both vertices a and b but does not 
contain the edge {a,b), and hence Mj is not an induced matching, contradiction. This completes 
the proof. ■ 

The triangle removal lemma of [21], which is one of the early major applications of the regularity 
lemma, asserts that for any e > there is a 5 = (5(e) > so that for > A^(e) any graph on N 
vertices from which one has to delete at least eA^^ edges to destroy all triangles contains at least SN"^ 
triangles. This and the above proposition implies that there are no (r, t)-RS graphs on vertices 
with r = Q{N) and t = Q{N). The original proof of [21] provides a rather poor quantitative relation 
between e and 5, but the improved recent proof of Fox [15] supplies better estimates (which are 
still very far from the known constructions). If the number of vertices is A^ and the graph is a 
pairwise disjoint union of t induced matching, each of size r = cN, then t is at most N/ log^^^ N, 
with X = 0(log(l/c)), where log^^^ A'^ denotes the x-fold iterated logarithm. For more details, see 
[15]. 
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3.2 Reconstruction Principle 

Here we prove lower bounds on the number of edges that a graph must miss, if it can be covered by 
disjoint induced matchings of size r. These lower bounds establish that the results in Section 2.2 are 
essentially tight for an important range of the parameters. Indeed, as proved in that section there 
are graphs on N vertices missing can be covered by disjoint, induced matchings 

of polynomial size. Yet, as we show below, any graph that can be covered by disjoint, induced 
matchings of size two or more must miss at least N^/"^ edges. We describe two proofs. The first is 
based on entropy considerations, and the second is an elementary combinatorial proof, that in fact 
yields a somewhat stronger result, as it bounds the minimum degree in the graph of missing edges. 
We believe, however, that both methods are interesting and each may have further applications. 
We start with the entropy proof. 

Let G = {V, E) be a graph on vertices that can be covered by disjoint induced matchings 
Ml, M2, ...Mt each of size r > 2. We will prove an upper bound on the number of edges \E\ based 
on an application of the reconstruction principle (and through information theoretic inequalities). 

To this end, we define a random variable A as follows: 

• Choose Mi uniformly at random 

• Choose an ordered set of two distinct edges ei , 62 from Mi 

Set A = (61,62). Let 61 = {W,X) and 62 = {Y,Z). Here we use upper-case letters to denote 
that each of these choices W, X, Y and Z is a random variable and we will use lower case letters to 
denote specific choices of these random variables. 

Claim 3.2. H{A) = log \E\ + log(r - 1) 

Proof: Since we choose each matching Mi uniformly at random, and each matching is of the same 
size (r), the first edge ei is chosen uniformly at random from the set \E\. Conditioned on the choice 
of 61, the remaining edge 62 is chosen uniformly at random from the r — 1 other edges in Mj. ■ 

Let dy be the number of missing edges incident to v £ V. Let be the set of non-neighbors 
of V, and let : D^, — )■ [d^] be a function mapping each non-neighbor of w to a unique integer in 
the set [dy]. 

• Choose A as above and let ei = {w, x) and 62 = (y, z) 

• Choose 5i with probability 1/2 to be either w or x, and let ^3 be the opposite choice 

• Choose 52 with probability 1/2 to be either y or z 

We set the random variable B = [si, /si(s2), /s2(s3)]- 
Lemma 3.3. H{B) > H{A) 

Proof: We prove that A can be computed as a deterministic function of B, and then we apply the 
Chain Rule for entropy to prove the Lemma. 

Claim 3.4. A can he computed as a deterministic function of B 
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Proof: Given B, we can compute S2 using si and fsi{s2), and using S2 and f 32(83) we can 
compute 53. This in turn defines the edge ei = (51,53) which uniquely determines Mj since the 
set of matchings disjointly covers the edges in G. From Mj and S2, we can compute the remaining 
edge 62- this is the unique edge incident to S2 in the matching Mj. ■ 

The Chain Rule for entropy yields the expansion H{B,A) = H{B) + H{A\B), but H{A\B) = 
because ^ is a deterministic function of B. We can alternatively expand H{B, A) as H{A)+H{B\A). 
Since H{B\A) > we get H{B) = H{B,A) > H{A), as desired. ■ 

Next, we give an upper bound for the entropy of B (based on the number of missing edges), 
and this combined with the Lemma above will imply a contradiction if the number of missing edges 
is too small. 

Definition 3.5. We will call a random variable S on V degree-uniform if S chooses a random 
vertex proportional to the degree in G. 

Claim 3.6. Si and S2 are degree-uniform random variables 

Note that these two random variables are not independent! 

Proof: We can choose the random variable A by choosing an edge uniformly at random from E, 
setting this edge to be ei and choosing €2 uniformly at random from the remaining edges in the 
matching Mi that contains ei. The distribution of Si in this sampling procedure (for A) is clearly 
degree-uniform . 

To prove the remainder of the Claim, we can slightly modify the sampling procedure for A. We 
could instead choose an edge uniformly at random from E and set this edge to be €2- Then choose 
an edge ei uniformly at random from the other edges in the matching Mj that contains €2- This is 
an equivalent sampling procedure for generating A, and from this procedure it is clear that 52 is 
degree-uniform. ■ 

Let d be the average degree in the complement of G. 
Lemma 3.7. H{B) < logiV + 21og(i 

Proof: We can decompose the random variable B into Bi = si, B2 = f 31(32) and -B3 = f 32(33). 
Again, using the Chain Rule for entropy we obtain that 



H(B) = H(Bi) + H(B2\Bi) + H(B3\B2, Bi) 



Since ^2 is a deterministic function of the random variables B2 and Bi , we get 



H(B:i\B2,Bi) = H(Bs\B2,Bi,S2) < H(Bs\S2) 



We can upper bound H(Bi) by logA^, and 



H(B2\Bi) = Y,Pr[Si = si]H(B2\Si = si) < J^Pr[5i = si]logd 




Using Claim 3.6, this is 



H(B2\Bi) = Y, 



N -l-d, 
2\E\ 



log dsi < ^ 



N-l-d 



2\E 



log d = log d, 
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where here we have used Jensen's Inequahty and the concavity of the functions logx and — xlogx. 
An identical bound holds also for H{B-:i\S2) again using Claim 3.6 and thus we get H{B) < 
log + 2 log J. ■ 

We can apply Lemma 3.3 and the bounds in Lemma 3.7 and Claim 3.2 to obtain the following 
theorem: 

Theorem 3.8. Let G = {V, E) he a graph on N vertices that can he covered hy disjoint induced 
matchings of size r > 2. Then the numher of missing edges satisfies 

We can apply a nearly identical argument in the case in which G is a bipartite graph: 

Theorem 3.9. Let G = {U,V,E) be a bipartite graph that can he covered hy disjoint induced 
matchings of size r > 3. Then the numher of missing edges satisfies 

\U\ X \V\ - \E\ > 0(r2/3|[/|2/3|I/|2/3), 

To prove this result, we choose A' to be three distinct edges from the matching Mj, and we 
use a length three path through pairs in U x V that are not in E to define the corresponding 
random variable B'. Again, the proof uses information theoretic inequalities and the fact that (if 
appropriately defined) A' can be reconstructed as a deterministic function of B' . It is worth noting 
that for a bipartite graph with \U\ = \ V\ = N and induced matchings of size 2, there is a simple 
construction missing only N edges. 

We can also give a direct counting argument, which is somewhat stronger, as it yields a lower 
bound on the minimum degree in the graph of missing edges. This counting argument proceeds by 
estimating the size of an appropriately defined set in two ways. Let G = {V, E) be an edge disjoint 
union of induced matchings Mi,M2, ...Mt each of size r. Again, let dy be the degree of v in the 
complement of G. 
Set 

3" = {{v, e)\v e V,e e E,v ^ e, 3Mj s.t. e G Mj and v is covered by MJ 
Lemma 3.10. |J| < min (^('^2") ' " 1 - - 1)) 

Proof: For each v £ V, v appears in precisely {N — 1 — dy){r — 1) elements of 3" since v belongs 
to exactly (A — 1 — d^) matchings and for each such matching there are exactly r — 1 choices of an 
edge (in the matching) that is not incident to v. 

Alternatively, each v G V also appears in at most ("^2") elements of 3": if {v, e) G 3" then V must 
not be a neighbor of each endpoint of e because the matching is induced. As there are at most ('^") 
choices of pairs of vertices that are not neighbors of v, the desired result follows. ■ 

We can also compute the size of 3" exactly: 

Lemma 3.11. |3"| = Ylvi'^ - 1)(^ - 1 - 4) 

Proof: For each edge e £ E, let Mi be the corresponding matching that covers e. There are exactly 
2(r — 1) choices for a vertex (covered by Mi) but not incident to e. Hence each edge e appears in 
exactly 2(r — 1) elements of 3". Thus 

m = 2(r -1)[(^]-^Y1 dv] = (r - 1) [n{N - 1) - 4] = J^(r - 1)(A - 1 - 4) 
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Combining the two estimates for |3~| we conclude that 



_ i)(Ar - l- d.) < J^min (f ^j, (A^ " 1 " 4)(r - 1)) 



V V 



It follows that for every v, the minimum term in the right hand side should be {N — 1 — d„)(r — 1), 
since otherwise the inequality cannot hold. Therefore we have proved the following, which implies 
Theorem 3.8. 

Theorem 3.12. If G = {V,E) is a graph on N vertices that is the disjoint union of induced 
matchings of size r, then the minimum degree d in the complement of G satisfies 



The assertion of Theorem 3.9 can be also proved by a counting argument. We omit the details. 

4 Applications 

4.1 Shared Communication Channels 

We apply our results to significantly improve the application in [10] of communicating over a shared 
directional multichannel. Roughly, when communicating over a shared channel we want the edges 
(corresponding to messages sent in some time step called a round) to form an induced matching. 
Otherwise, a receiver will hear messages sent from two different sources and the messages will 
appear garbled. Birk, Linial and Meshulam construct graphs with positive density that can be 
covered by roughly induced matchings where r = (log A^)^('°s'°s^/('°s^°s'°s^) ). The authors 
then use these graphs to design a communication protocol for N stations over a shared directional 
multi-channel where the round complexity of this protocol is O(^). This is a slightly better than 
poly-logarithmic improvement over the naive protocol for bus-based architectures. 

We can use our constructions to achieve a round complexity of 0{N'^~^) over a shared directional 
multi-channel. This is the first such protocol that provides a polynomial improvement over the 
naive protocol. We accomplish this using just one transmitter and two receivers per station. This 
corresponds to a partition of the edges of a complete bipartite graph into two graphs each of which 
can be decomposed into a small number of induced matchings. If we allow C = C(e) receivers 
per station, we can achieve a round complexity that is 0{N'^~^'') for any e > (here iV is a trivial 
lower bound). Hence, while previous protocols required a nearly quadratic number of rounds with 
a constant number of receivers per station, our protocols require only a nearly-linear number of 
rounds. 

Motivated by the application to communication over a shared channel, Meshulam [2(1] conjec- 
tured that any graph on vertices with positive density cannot be covered by 0{N'^~^) induced 
matchings. The constructions presented in Section 2.1 and in Section 2.2 disprove this conjecture 
in a strong sense. 

First we explain the model considered in [10]. Roughly, the goal is to design a good communi- 
cation protocol using a small number of shared communication channels. More precisely, suppose 
we have N stations, and each wants to send a (distinct) message to every other station. We fur- 
ther assume that each message is (roughly) the same size. In this context, it is often prohibitively 
expensive to build a point-to-point communication channel from each station to every other one. 
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Often, the proposed solution is to use some form of a shared communication channel. Indeed, the 
standard bus-based architecture connects all pairs of stations using a single connection in such a 
way that only one message can be sent on the channel per time step and hence a total of N'^ rounds 
are needed to send all messages. 

There are other architectures that can be implemented cheaply in hardware and can accomplish 
this task in a smaller number of rounds. One such architecture is the shared directional multi- 
channel. The combinatorial abstraction is that we imagine the communication graph as a complete 
bipartite graph Kjsf^jsf (with vertices on the left, representing the transmitters of the stations, 
and N vertices on the right, representing the receivers). A directed multi-channel allows us to 
partition Kn,n into C graphs Gi,G2, ...Gc- These graphs correspond to allocating c receivers to 
each station. For each graph Gj, in each round we can exchange all messages corresponding to 
the edges in some induced matching in Gi in one time step. These matchings are required to be 
induced because otherwise messages would interfere in the underlying hardware. 

Thus the problem of designing a communication protocol for this architecture that completes 
in a small number of rounds and does not use too many transmitters and receivers per station is 
exactly the problem of covering all the edges of a complete bipartite graph (using at most C graphs) 
so that the number of induced matchings needed to cover the edges in each graph is small. The 
number C represents the number of receivers that each station must be equipped with, assuming 
it has only one transmitter, and so our goal is not only to minimize the number of rounds, but also 
to do so for a small value of C. 

• For C = 2, we give a protocol that completes in 0{N'^~^) rounds for 5 > 0.058 and 

• For any e > 0, we show that there is a C = C(e) = 2'^^^^ so that there is a communication 
protocol that completes in 0{N^~^'') rounds 

Let Kn^i^ be the complete bipartite graph with vertices on the left and A^ on the right. 

Theorem 4.1. There is a partition of the edges of Ki^^i\j into two graphs Gi and G2 so that each 
of these graphs can be covered by at most 0{N'^^^) induced matchings, for 6 > 0.058. 

Proof: This follows immediately from the construction given at the very end of Section 2.2: we 
can choose G'l and G2 that cover all edges of K^^n, where G'l covers all edges of Kj\f^]y but at most 
N'^~^ and yet it is a union of at most N'^~^ induced matchings. The second graph G2 consists of 
all these remaining edges. Since it contains at most N"^"^ edges in total we can cover G2 by trivial 
induced matchings - one for each edge in The total number of induced matchings in each graph 
is thus at most OiN"^'^). ■ 

Theorem 4.2. For any e > 0, there is a G = C(e) = 2'^(7) so that the edges of Kn, N can be 
partitioned into Gi,G2, --Gc and each of these graphs can be covered by at most 0{N^^'^) induced 
matchings. 

Proof: To obtain this theorem, we can instead invoke the construction in Section 2.1 to obtain a 
bipartite graph G (obtained by splitting the vertices of the graph constructed in that section into 
two equal parts and by keeping all edges that join vertices in the two parts). For each i, we can 
take Gi to be a random shift of G - i.e. we construct Gi by permuting the labels of the vertices 
on the right randomly. G is missing less than A^^""^ edges, for 6 = 2~^^~^\ and hence if we take 
C = 2/(5 random shifts the expected number of edges that are not covered in any Gi is less than 
one. Hence there is some choice of Gi,G2, ■■■Gc that covers the edges in the complete bipartite 
graph and yet the edges in each Gi can be covered by at most 0{N^~^'^) induced matchings. We 
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note that the above proof can be derandomized using the method of conditional expectations, that 
is, the graphs Gi can be generated efficiently and deterministically. ■ 

Finally we mention a simple lower bound for the number of rounds needed, proved by Meshulam 
[20] . This shows that for any constant number of receivers a super-linear number of rounds is needed: 

Proposition 4.3 ([20]). For any partition of the edges of K^^^j^ into Gi, G2, ■■■Gc, the total number 
of induced matchings needed to cover Gi, G2, ■■■Gc is at least b{G)N^~^^/^'^'^ . 

Proof: We apply induction on C, the result for C = 1 is trivial. Consider the case C = 2. Without 
loss of generality, let Gi contain at least half of the edges from the complete bipartite graph and 
suppose that the minimum number of induced matchings needed to cover Gi is A^''. Then there 
is an induced matching (in this set) that contains at least ^N"^'^ edges and hence G2 contains 
a complete bipartite graph where the number of vertices on the left and on the right is at least 
ijy2~r^ Hence the number of induced matchings needed to cover G2 is at least j^A^^~^^. Since the 
quantity max(A^'', A^^"^'') is minimized for r = 4/3 the total number of induced matchings needed 
to cover Gi and G2 is at least n{N'^/^). 

We can iterate the above argument in the general case. Without loss of generality let Gi 
contain at least ^N"^ edges and suppose the minimum number of induced matchings needed to 
cover Gi is N'^ . Then the union of G2,G^, ■■■Gc contains a complete bipartite graph where the 
number of vertices on the left and on the right is at least ^N'^~^ . We can assume by induc- 
tion that the total number of induced matchings needed to cover G2,G^, ■■■Gc is at least some 
6'(C)7V(2-r)(i+i/(2c^-i-i))_ Tj^g quantity 

2C-1 

max(r, (2 - r) x _ ^ ) 

is minimized for r = and this completes the proof. ■ 

Hence any protocol requires at least ^^(log i) receivers per station to reduce the number of 
rounds to 0{N^~^'')^ In contrast, the protocol in Theorem 4.2 uses receivers per station to 

complete this same task in 0{N^~^'') rounds. 

4.2 Linearity Testing 

Here we observe that our graphs can be plugged in the analysis of Hastad and Wigderson [19] of 
the graph test of Samorodnitsky and Trevisan [22] to provide a (modest) strengthening. We obtain 
slightly better bounds on the soundness of this test, which may be of interest for a particular range 
of the parameters. 

The classical linearity test of Blum, Luby and Rubinfeld chooses a pair of points x and y 
uniformly at random from the domain of a function, and checks if f{x) + f{y) = f{x + y). The test 
accepts / if and only if this condition is met, and indeed this test always accepts a linear function 
and if / is not linear, the probability that this test accepts / can be bounded by | + -y^, where 
d{f) is the maximum correlation of / with a linear function [11]. 

What if we want to reduce the probability that a function / that is not linear passes this test? 
We could perform r independent trials, in which case the probability that / is accepted is bounded 
by (i + However such a test queries the function / on 3r locations. Motivated by the 

problem of designing a PCP with optimal amortized query complexity and the related problem for 
linearity testing, Samorodnitsky and Trevisan introduced a graph-based linearity test: Associate 
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each vertex in an r-vertex complete graph with a randomly chosen element from the domain of /, 
and for each edge check if f{x) + f{y) = f{x + y) where x and y are the values associated with the 
endpoints of the edge. This test accepts if and only if all of these conditions are met. 

This test queries the function / on r + (0 locations and the hope is that the soundness should 
behave approximately like (2) independent trials of the original linearity test [11]. Samorodnitsky 
and Trevisan [22] showed that the soundness of this test is bounded by 

This analysis was subsequently simplified and improved by Hastad and Wigderson [19] - using 
the known existence of graphs that have many edges but can be covered by large (disjoint) induced 
matchings. The intuition behind this connection is that an induced matching corresponds to 
independent trials of the original Blum-Luby-Rubinfeld linearity test (althoug the formal analysis 
somewhat masks this intuition). Hastad and Wigderson [19] proved: 

Theorem 4.4. If G = {V,E) is an {r,t)-RS graph, then the graph-test for G accepts a function f 
with probability at most 

e-'-*/8 + d{fY/\ 

Hastad and Wigderson [19] used the construction of Ruzsa and Szemeredi [21] mentioned in 
the introduction, which shows that there are (r, t)-RS graphs on N vertices with r = o{viosn) ^^'^ 
t = N/3. 

We can plug our constructions directly into this theorem to obtain slightly better bounds, for 
some special values of d{f). Our constructions are dense, and hence improve the first term in the 
bound, but the second term is slightly worse (although we still have r = A^^~°^^^). In general, the 
obtained bounds will be better than either of those in [22] or [19] for some values of d{f). Note 
that as the complete graph on vertices contains every graph on N vertices, these bounds, like 
the ones of [22] and [19], provide an upper estimate for the probability that the complete graph 
linearity test on N vertices accepts a function /, showing that it is at most 

min( 2-(") +ci(/),2-^^-°'^' +d(/)^^-°<^2"^(^^) +rf(/)^^-°''^' ). 

The first term in the minimum is the bound of [22], the second is that of [19], and the third (in 
which the o'(l) term is a bit worse than the one in the second) follows from our graphs. 

4.3 The Directed Steiner Tree Problem 

In this short subsection we briefly note the connection between our constructions and a candidate 
randomized rounding algorithm for the directed Steiner tree problem that motivated Vempala [25] 
to ask about the existence of certain (r, t)-RS graphs. 

Giving a poly-logarithmic approximation algorithm for the directed Steiner tree problem is a 
famous open problem in approximation algorithms. A special case is the group Steiner tree problem 
(in an undirected graph), for which Garg, Konjevod and Ravi gave an elegant, poly-logarithmic 
approximation algorithm [IS]. Charikar et al [12] give an approximation algorithm for the directed 
Steiner tree problem whose approximation guarantee is 0{N^) for any e > 0, and this guarantee 
can be made poly-logarithmic at the cost of running in quasi-polynomial time. 

Even our understanding of the naive linear programming relaxation is quite weak. Zosin and 
Khuller [26] give a i}{^/k) integrality gap (where k is the number of terminals), but this construction 
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has exponentially many (in k) vertices. Hence we could still hope that the naive relaxation has at 
most a poly-logarithmic (in A^) integrality gap. 

Rajaraman and Vempala considered a stronger relaxation and a candidate rounding algorithm. 
In the case in which the support of the solution to the linear program is a tree, they proved that their 
rounding algorithm achieves a poly-logarithmic approximation ratio and this analysis is reminiscent 
of the rounding procedure for the group Steiner tree problem [In]. 

However, even when flow merges in one layer of a layered graph (i.e. when the fractional 
solution is not supported on a tree), attempting to analyze the behavior of the rounding algorithm 
led Vempala to a combinatorial conjecture: 

Conjecture 1. /,:' ">/ Let G = ([/, V, E) he an N y. k complete bipartite graph and N > k. Let CP be 
a partition of the edge set and for a part p £ T, let pi denote the degree of vertex i in p (i.e. the 
number of edges of p incident to i). Then 

y minfl,y^)>C^. 

iGU,j&V p& *^ 

Our constructions yield a negative answer to the above conjecture. In our negative example 
we have N = k. To obtain this result, we can instead invoke the construction in Section 2.2 to 
obtain a bipartite graph H (obtained by duplicating the vertices of the graph constructed in that 
section). We can take "P to be the induced matchings covering H and additionally we add a part 
in the partition (consisting of a single edge) for each edge across the bipartition missing from H. 

We can upper bound the right hand side as: 

PiPj\ ^ V- v-KPi 



H is an (r,t)-RS graph (and r = VL{N'^-f) in our construction) and so for each part p we have 
S(i j)G// ~ ^ because p is an induced matching with respect to H. The number of parts in 
the partition (ignoring singletons, which are not in H anyways) is at most 0{N^) and so we can 
bound the contribution of the first term by 0{Nf). Also, the number of edges that H is missing 
(across the bipartition) is at most 0{N'^) and hence we can bound the above sum by 0{N'' + Nf) 
for e and / as in Theorem 2.21 (recall that N = k). Since we can have both e and / at most 1.942 
it follows that the conjecture is false. 



5 Concluding Remarks and Open Questions 

We have given two constructions of nearly complete graphs that can be decomposed into large 
pairwise edge disjoint induced matchings and described several applications of these graphs. 

The main combinatorial open problem that remains is to determine or estimate more precisely 
the set of all pairs (r, t) so that there are (r, t)-RS graphs on vertices. This is interesting for 
most values of the parameters, but is of special interest in some specific range. In particular, if 
for r = (io^-)3 with g > 1, one can show that t = o{N), this will improve the best known upper 
bound for the maximum possible cardinality of a subset of {1, 2, . . . , N} with no 3-term arithmetic 
progressions-a problem that received a considerable amount of attention over the years (see [2.3] 
and its references). 

The study of the combinatorial problem above seems to require a variety of techniques: the 
known constructions of [21], [14], [20] and the ones given here apply tools from Additive Number 
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Theory, Coding Theory, low degree representations of Boolean functions and Geometry, while the 
proofs of non-existence rely on the regularity lemma and on combinatorial and entropy based 
techniques. All of these, however, still leave a wide gap between the upper and lower bounds for 
at least some of the range, and it will be interesting to find additional ideas that will help to study 
this problem. 

In all the applications considered here there are still remaining open problems. The communi- 
cation protocols over a shared directional multi-channel we suggest, while improving substantially 
the existing ones, are still not optimal, and the problem of deciding the best possible number of 
rounds for stations, even with two receivers per station, is still not settled, although our results 
show that it is N"^"^ for some S between 0.058 and 2/3. The best possible upper bound for the 
probability of acceptance of a function / in the linearity graph test, using a complete graph of size 
A^, is also not precisely determined as a function of A^ and d{f) (although here the gap between the 
upper bounds and the lower bounds is not large-see [19].) Finally, it will be interesting to decide 
if our graphs can be helpful in establishing new integrality gap results for the natural relaxation 
of the directed Steiner tree problem, rather than merely estimating the performance of specific 
rounding schemes. 
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